Skip to content

Fix result handoff and deferred action continuations - #3073

Merged
George Ng (GeorgeNgMsft) merged 4 commits into
mainfrom
georgengmsft-result-entity-handoff-fixes
Sep 25, 2026
Merged

George Ng (GeorgeNgMsft) merged 4 commits into
mainfrom
georgengmsft-result-entity-handoff-fixes

Conversation

@GeorgeNgMsft

@GeorgeNgMsft George Ng (GeorgeNgMsft) commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

This standalone product fix addresses intermediate-result handoff failures surfaced by the GHCP evaluation. Successful actions no longer fail simply because translation assigned a result label without an entity, deferred requests remain in the execution queue, and concrete result references are resolved before their consumers run. The PR targets main directly, without eval-harness changes or a dependency on #3072.

  • Separate optional entity-name references, explicit resultValue references, and deferred translation context. Missing consumed data still fails; display text and structured display rawData are not silently substituted as typed values.
  • Preserve request-local completed-action snapshots across deferred continuations, even without conversation history/memory extraction. Translate only the remaining request without replaying producers or automatically switching to reasoning.
  • Restore concrete-value lookup and strictly validate substituted values against the consumer schema. Preserve empty values and treat reference-looking returned strings as data. Reject undeclared/duplicate references.
  • Stop remaining legacy actions while a user choice is pending, including independent and returned additional actions. Keep the choice available, explicitly report that remaining steps will not resume automatically, and disable reasoning fallback. Preserve standalone choices and structured execution's awaited choice handling.
  • Bound deferred context to a 64 KiB serialized UTF-8 envelope containing the remaining request and full history. Project result data without execution metadata, display alternates, or identical representations; retain distinct display data even when history text is only a summary. Oversized context stops before continuation translation without truncation, summarization, or replay. This is not a total model-token budget; schema prompts are separate. Concrete $result substitution remains lossless and outside this limit.
  • Preserve the caller's active-schema/schema-family scope during deferred translation. Recheck execution eligibility before enqueueing translated continuations; unavailable scope, unknown actions, and execution-disabled actions cannot silently continue or trigger producer replay.
  • Document the contracts and add 29 handoff cases, 17 bounded-context cases, 7 deferred-scope cases, and 11 schema-validation cases.

Evidence and scope

The recorded evaluation contains nine explicit missing-result-entity errors across readFile, getList, clearList, prFiles, and prChecks, each after handler success. One clearList mutation persisted before the failure. Full translated plans were not retained, so the traces do not establish which labels were unused versus consumed. This PR fixes the eager invariant and independently reproduced queue/resolver defects without inventing handler entities or claiming all nine workflows now complete. No live eval was rerun and no success-rate improvement is claimed.

Automatic queue resumption after a legacy choice and large-output retrieval/chunking are intentionally not implemented.

Validation and review

  • The original product patch was isolated onto main without changing its contents; eval commits were removed. The original isolated head passed the full dispatcher suite (135 suites / 2,152 tests, 12 existing skips) and action-schema suite (4 suites / 332 tests, 62 existing skips).
  • Two independent adversarial reviews examined base 5d5e23fa6 through bounded-fix head 725a624a6. Two newly exposed deferred-path issues were independently validated: missing execution-eligibility checks and dropped caller schema restrictions. Both were then corrected in 2029d4fc1; six regression failures were reproduced before the corrections. Review comments were not posted or resolved.
  • The latest dependency-aware build and 136 targeted tests across five suites pass, covering handoffs, size boundaries, scope restrictions, structured execution, and chained choices. Pinned formatting and all four ratchets (lint, complexity, circular dependencies, test debt) pass against origin/main.
  • Size tests cover exactly 64 KiB, one byte over, UTF-8 encoding, aggregate/inherited context, empty values, and no producer replay. Scope tests stop at an offline translator boundary; no model calls are needed.
  • Hosted verification for 2029d4fc1029cf4e9e34872a076f5943c88f5983: build-ts run 36187276971 passed all six Windows/Linux/macOS × Node 22/24 jobs. All applicable PR checks also pass, including Linux/Windows Shell & CLI smoke tests, shell packaging on all three operating systems, .NET, CodeQL, formatting, repository policy, and documentation generation. The fork-only formatting check is skipped as expected.

The original Linux/macOS CI failures were inherited eval portability issues, fixed separately in #3072 (b47f8bacd), whose six-job OS/Node matrix is green. #3072 and #3077 remain in their own eval stack.

Comment thread ts/packages/dispatcher/dispatcher/src/execute/actionHandlers.ts Outdated
Comment thread ts/packages/dispatcher/dispatcher/src/translation/pendingRequest.ts
Separate optional entity names from concrete result values and deferred translation context. Preserve completed outputs without replaying producers, reject missing or invalid consumers, and add offline regressions.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@GeorgeNgMsft
George Ng (GeorgeNgMsft) force-pushed the georgengmsft-result-entity-handoff-fixes branch from 964c682 to 73d9e68 Compare September 25, 2026 19:21
@GeorgeNgMsft
George Ng (GeorgeNgMsft) removed this pull request from stack #3074 September 25, 2026 19:26
@GeorgeNgMsft
George Ng (GeorgeNgMsft) changed the base branch from georgengmsft-ghcp-eval-implementation to main September 25, 2026 19:26
Report unexecuted remaining steps without automatic continuation, reasoning fallback, or replay. Preserve standalone choices and structured choice handling.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Project result data without execution metadata or duplicate representations and enforce a 64 KiB serialized UTF-8 context limit before deferred translation. Preserve distinct data and fail explicitly instead of replaying producers or translating from truncated results.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Recheck translated continuations before enqueueing effects and preserve active schema/family restrictions. Stop explicitly without replay when scope is unavailable or actions are disabled.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@GeorgeNgMsft
George Ng (GeorgeNgMsft) added this pull request to the merge queue Sep 25, 2026
Merged via the queue into main with commit fff27a2 Sep 25, 2026
27 checks passed
Dominic Nguyen (datduyng) pushed a commit to datduyng/TypeAgent-1 that referenced this pull request Oct 1, 2026
…ity case (microsoft#3077)

This follow-up above microsoft#3072 adds bounded safe-read recovery, frozen-run
checks and scoped file-handler consent, integrated with the protocol-5
common-file corpus and Luna-pinned evaluation package. It preserves
seven candidate semantics and four five-case cohorts without rewriting
historical measurements. The independent result-handoff fixes in microsoft#3073
are not a dependency.

- Integrate parent `f2f045c20` through a normal merge, preserving bot
history. `common-files-v1` replaces nine list-dependent cases with
common-capability file tasks: all 20 cases apply to every candidate,
yielding 140 executions per full repetition. This supersedes protocol
4's 131 executions plus nine native N/A slots, not historical scores.
- Derive counts from the frozen schedule and reject changed protocols,
corpus versions or orders. Require fresh corpus-marked preflight
evidence; preserve file snapshots, byte-exact backups, unrelated-state
checks and failed-trial evidence.
- Retain narrow per-case permissions, clarification gates and M3's
original-backup prerequisite. No arbitrary shell-write approvals or
measured list mutations. Preserve historical A4's original prompt,
failure and replacement rationale without claiming general product
ambiguity is solved.
- Recover only from affirmative read-I/O evidence plus a complete
read-only backend trace for TypeAgent. Write/copy failures, denials,
cancellations, uncertainty, missing details and conflicting error fields
remain terminal. No retry loop or expanded artifact access.
- Handle the file handler's second Run/Cancel consent after dispatcher
confirmation. Require the exact prompt/choices, unambiguous current-tool
admitted action and independent case oracle. Bind structured consent to
scope/operation/interaction, consume once, invalidate on
unrelated/terminal work, and compare response fields independent of JSON
key order. Consent never supplies an ambiguous referent.
- Preserve the sibling evaluation package, short plugin README link,
initial-ballpark/environment-isolation caveats, Luna pin, early ledger
admission and ledger-derived limits.

**Validation:** dependency-inclusive build and 137 dispatcher tests
passed during integration. Final consent corrections pass 34 evaluation
tests, package syntax build, pinned formatting and all four
committed-head ratchets against `79fad5b4f`: lint 0→0,
cyclomatic/cognitive over-budget counts 0→0, circular dependencies
220→220 with 11 existing exceptions, test debt 0→0. Fresh preflight
completed actual inventory/read/copy/append and read-only GitHub/network
operations. Structured S4 pilot verified both consent stages and exact
state preservation. Other pilot denials remain failures; pilot evidence
is separate from measured outcomes. Hosted CI for the latest head is not
yet claimed complete.

**Fresh live round — paused after 21/140 trial records:** the user
separately authorized 40,000 additional Copilot AI credits. Execution is
frozen at `96e904546acc389cf64284db9b6fbcbe57191344`. Three complete
seven-candidate batches ran for M5, M1 and M4; 119 trials remain
unstarted. The runner paused before batch four because a timed-out M4/C3
request lacks billing settlement. Its full 250-credit reservation
remains held; no automatic retry or trial replay occurred. This is a
reconciliation pause, not budget exhaustion. Partial evidence and
accounting snapshots are preserved privately. Independent
faithfulness/accuracy review is outstanding; no new accuracy or latency
conclusion is claimed.

Historical ledgers remain closed and the protocol-2 run at `e847c7a009`
remains unchanged. Private network payloads, artifact paths and billing
traces are not published. Stack #3080 remains microsoft#3072 followed by this PR;
no stack metadata changes.

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: typeagent-bot <typeagent-bot[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants